Papers with multimodal research

4 papers
FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation (2025.acl-demo)

Copied to clipboard

Challenge: FlagEvalMM is an evaluation framework designed to assess multimodal models . it is designed to be used for vision-language understanding and generation tasks .
Approach: They propose an evaluation framework that decouples model inference from evaluation through an independent evaluation service.
Outcome: The evaluation framework offers accurate and efficient insights into model strengths and limitations.
Quantifying the Visual Concreteness of Words and Topics in Multimodal Datasets (N18-1)

Copied to clipboard

Challenge: Existing work suggests that concepts with concrete visual manifestations are easier to learn than abstract ones.
Approach: They propose an algorithm for automatically computing the visual concreteness of words and topics within multimodal datasets.
Outcome: The proposed algorithm predicts the capacity of machine learning algorithms to learn textual/visual relationships.
Image Position Prediction in Multimodal Documents (2020.lrec-1)

Copied to clipboard

Challenge: Existing multimodal tasks allow machines to understand images by describing or being asked in natural language.
Approach: They propose a task that predicts the positions of images in a given document . they use a dataset of 66K multimodal documents with 320K images from Wikipedia .
Outcome: The proposed task outperforms baselines while the performance is far from human.
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in multimodal large language models have led to progress in tackling complex reasoning tasks that combine textual and visual information.
Approach: They introduce a robust version of the Massive Multi-discipline Multimodal Understanding and Reasoning (MMMU) benchmark.
Outcome: The proposed model performs lower on MMMU-Pro than on the previous benchmark, ranging from 16.8% to 26.9%.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations